You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Adds support for SQL Server's vector(N, float16) base type, opted into by a new connection string keyword.
System.Half exists on .NET but not on .NET Framework, so the two differ by necessity:
.NET — a float16 column is read and written as SqlVector<Half>, exchanged exactly as the server sent it.
.NET Framework — there is no native representation, so the column is surfaced as a JSON string, and can also be read as SqlVector<float>, which widens losslessly. Writes go back as a JSON string or a SqlVector<float>.
Opting in
Vector Type Support (also accepted as VectorTypeSupport) takes off, v1 or v2, and defaults to v1 — the same name, values and default as the JDBC driver's vectorTypeSupport. Upgrading the driver therefore does not change the representation an existing application receives; float16 arrives in binary form only when the application asks for v2.
A base type can only be used once the connection has negotiated the version covering it. The driver rejects a value above that version itself, naming the keyword to set, rather than leaving it to the server — which was measured to accept it silently at v1 and to report a malformed protocol stream at off. The acknowledgement is likewise refused if it exceeds the version requested.
Conversion between base types
Whether SQL Server converts between float32 and float16depends on the server: some builds block it and report error 42238. Verified as working against SQL Server vNext CTP 1.0 (18.0.258.0). A JSON string is accepted by every server, so it is the portable way to write a column whose base type differs from the value's — and on .NET Framework the only way where conversion is blocked. Tests which depend on the conversion skip rather than fail.
Bulk copy
The INSERT BULK statement states the destination column's base type, and the server then requires a binary payload of exactly that width and performs no conversion within the data stream. Measured at every negotiated version: a text declaration for a vector column is refused with Invalid column type from bcp client, and text sent under a vector declaration is refused with a length mismatch.
So any in-memory value — a JSON string, or a SqlVector<T> whose element type differs from the column's — is converted to the destination's base type by the driver, with a range check, which does not depend on the server. A payload read from another vector column keeps its own base type, so copying between columns of different base types is reported by the server rather than silently narrowed. Both match the JDBC driver.
A connection below v2 is told a float16 column is a varchar(max), so bulk copy into one still works at the default level: the value travels as text and the server converts it.
Public API
One addition: the SqlVectorTypeSupport enum and SqlConnectionStringBuilder.VectorTypeSupport. SqlVector<T> is unchanged.
Tests
float16 added to the existing generic native vector suite, run over a v2 connection
the same suite run through SqlVector<float> on every framework, which is the .NET Framework representation and where the hand written binary16 codec is the production path
version negotiation per keyword value, the client's own version ceiling, an acknowledgement above the requested version, and rejection of a base type below it
the binary16 codec validated against System.Half across all 65,536 binary16 patterns, and across the binary32 space by sampling (a full sweep of all 4,294,967,296 patterns was run once during development and reported no divergence, but is too slow to repeat on every CI leg)
bulk copy from textual, object-typed, in-memory vector and raw-payload sources, including a column wider than the float32 element limit
rounding parity at binary16 ties between bulk copy and a server-parsed INSERT
a DbDataAdapter round trip of a float16 column on .NET Framework
A vector column's base type and number of dimensions were already available
from the column schema, but only as a numeric scale and a column size which
the caller had to decode. They are now surfaced under their own names, so
that applications inspecting result set metadata do not have to know that
encoding.
Also registers the vector type in the DataTypes schema collection, where it
was missing.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: e17ed782-c576-4cb7-9b4b-7ad286d7a7d0
SQL Server transports vector(N, float16) elements as raw binary16 values.
System.Half is only available on .NET, so the conversion is implemented
manually for .NET Framework.
The manual implementation is compiled for every target framework rather than
only for .NET Framework, so that it can be validated exhaustively against
System.Half on .NET while remaining the code path .NET Framework actually
uses. It is verified against every binary16 bit pattern, a strided sweep of
the single precision range, and the rounding, subnormal, overflow and
underflow boundaries.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: e17ed782-c576-4cb7-9b4b-7ad286d7a7d0
Advertises version 2 of the VECTORSUPPORT feature extension, so that a
vector(N, float16) column is exchanged in its native binary form rather than
as a varchar(max) JSON string.
On .NET such a column is surfaced as SqlVector<Half>. .NET Framework has no
System.Half, so it is reported as a string there, matching how it is already
presented when the server does not negotiate float16 support. Callers on
either framework can explicitly request a strongly typed value via
GetSqlVector<float>, which widens the elements without loss.
SqlVector<T> continues to derive the base type written to the wire from T
alone. Conversion between base types is left to the server, which performs it
for parameters. Bulk copy is the exception: it declares the destination's base
type in the INSERT BULK statement, so a payload using a different base type is
rejected as a column length error rather than converted, and is rewritten by
the driver first. That conversion runs after coercion, because the payload
coercion produces uses the source value's own base type: a JSON string always
yields float32, which is how a float16 column reads back where System.Half is
unavailable.
SqlVector<T>.ToString() now returns the vector's values as a JSON array rather
than the type name.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: e17ed782-c576-4cb7-9b4b-7ad286d7a7d0
Describes the base types a vector column can have, how they map to
SqlVector<T>, and how a float16 column is read and written on .NET Framework,
where System.Half does not exist. Also documents the vector feature extension
versions and the column metadata properties.
Adds a sample covering both frameworks, reading a float16 column as an exact,
widened or JSON value, inspecting a column's base type and dimensions, and
converting between base types.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: e17ed782-c576-4cb7-9b4b-7ad286d7a7d0
The reason will be displayed to describe this comment to others. Learn more.
Pull request overview
Adds end-to-end support for SQL Server vector(N, float16) by negotiating VECTORSUPPORT feature extension version 2, introducing an IEEE-754 binary16 codec, and wiring float16 handling through SqlVector<T> read/write paths (including bulk copy), with accompanying docs and tests.
Changes:
Negotiate vector feature extension v2 and track negotiated vector capability version (float32/float16) on the connection.
Add float16 vector support across SqlVector<T>, SqlDataReader, SqlBuffer, SqlParameter, SqlCommand, and SqlBulkCopy, including payload conversion for bulk copy.
Add unit/manual tests plus docs/snippets/sample updates; expose vector base type + dimensions via DbColumn indexer and register vector in the DataTypes schema collection.
Reviewed changes
Copilot reviewed 23 out of 23 changed files in this pull request and generated 2 comments.
ConvertPayloadElementType doesn’t validate the vector header magic/version bytes before using the length and element type fields. This can cause non-vector payloads to be converted (or to fail later with less appropriate exceptions). Validate VecHeaderMagicNo/VecVersionNo up front, consistent with GetCountsOrThrow.
if (tdsBytes.Length < TdsEnums.VECTOR_HEADER_SIZE)
{
throw ADP.InvalidVectorHeader();
}
Bulk copy read a vector column through the representation the reader
surfaces, which is a JSON string on frameworks without System.Half. That
round trip is both larger than the payload it encodes and unable to carry a
negative zero, because System.Text.Json on .NET Framework serialises one as
zero and parses a negative zero literal back as positive zero.
Reading the payload directly avoids both. It is chosen once per column, when
the source and destination are both vector columns, alongside the existing
decimal and streaming decisions. Any difference in base type between the two
is still resolved when the value is converted.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: e17ed782-c576-4cb7-9b4b-7ad286d7a7d0
SqlTypes.SqlVector<float> doesn’t resolve to any namespace/type in this file (there’s no using SqlTypes = ... and no SqlTypes namespace). This should be fully qualified to Microsoft.Data.SqlTypes.SqlVector<float> (or add an alias) to avoid a compile error.
// The payload is converted directly rather than through a strongly typed vector,
// so that .NET Framework, which has no System.Half, can also write to float16
// destinations.
return SqlTypes.SqlVector<float>.ConvertPayloadElementType(payload, destinationElementType);
FromTdsPayload reads header fields (element type/length) without validating the vector magic/version bytes. This makes the widening path accept malformed payloads that GetCountsOrThrow would reject.
if (tdsBytes.Length < TdsEnums.VECTOR_HEADER_SIZE)
{
throw ADP.InvalidVectorHeader();
}
ConvertPayloadElementType should validate the vector header magic/version before interpreting element type and length; otherwise malformed byte[] values can be converted and sent on the wire rather than failing fast with InvalidVectorHeader.
if (tdsBytes.Length < TdsEnums.VECTOR_HEADER_SIZE)
{
throw ADP.InvalidVectorHeader();
}
doc/samples/SqlVectorFloat16Example.cs:146
These interpolated strings won’t compile because the expression uses double quotes (e.g., column["VectorBaseType"]) inside a double-quoted string literal. Escape the quotes (or assign to a local variable) before interpolating.
Console.WriteLine($"\nColumn base type: {column["VectorBaseType"]}");
Console.WriteLine($"Column dimensions: {column["VectorDimensions"]}");
The existing suite covers nulls where the source and destination share a base
type, but not where they differ, which is the path that converts the payload.
Verified that nulls survive in every combination, interleaved with non-null
rows so that a row's nullness cannot be satisfied by position alone.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: e17ed782-c576-4cb7-9b4b-7ad286d7a7d0
The reason will be displayed to describe this comment to others. Learn more.
Pull request overview
Copilot reviewed 23 out of 23 changed files in this pull request and generated no new comments.
Suppressed comments (2)
doc/samples/SqlVectorFloat16Example.cs:146
These interpolated strings won’t compile because the expression contains a string literal with double quotes (e.g., column["VectorBaseType"]) which terminates the outer interpolated string. Assign the indexer results to variables (or constants) first, then interpolate those variables.
Console.WriteLine($"\nColumn base type: {column["VectorBaseType"]}");
Console.WriteLine($"Column dimensions: {column["VectorDimensions"]}");
This SqlVector special-case is redundant/unreachable because SqlVector implements ISqlVector (so it will already be handled by the earlier value is ISqlVector branch). Keeping the extra branch increases maintenance burden and risks diverging behavior.
else if (currentType == typeof(SqlVector<Half>))
{
value = ((ISqlVector)value).VectorPayload;
}
#endif
The reason will be displayed to describe this comment to others. Learn more.
Summary
This adds vector(N, float16) by advertising VECTORSUPPORT v2, introducing a hand-written binary16 codec, teaching SqlVector<T> about System.Half, and switching vector→vector bulk copy to a raw-payload transfer. The engineering is careful, and I want to call out specifically that the endianness and element-size arithmetic is correct throughout — I went looking for a missed (ColumnSize - 8) / 4 and there isn't one. The commit split is clean and the rationale in the description is unusually good.
I have one blocking correctness issue, plus a set of suggestions. Inline comments carry the detail and suggested diffs; this is the map.
Blocking
SqlBuffer.GetSqlVector<T>() succeeds or throws depending on the row's nullness. The IsNull branch builds a vector for anyT without consulting the column's base type, while the non-null branch validates. GetSqlVector<Half>() over a float32 column therefore returns Null for NULL rows and throws NotSupportedException for non-NULL rows in the same result set. Data-dependent rather than schema-dependent, so it is hard to find in testing and impossible to guard against in caller code.
Suggestions
The codec is not bit-exact against System.Half for NaN, contrary to the description, and the tests are written to step around exactly that case (continue / IsNaN-only). float.NaN carries the sign bit, so widening flips the sign of every NaN relative to System.Half; narrowing canonicalises payloads that (Half)float preserves. Either match System.Half or pin the canonical form with explicit assertions and drop the "bit-exact" claim.
Bulk copy's vector case uses metadata.scale rather than the scale local that the surrounding code establishes for encrypted columns.
The widening path skips the magic-number and version validation that the matching path gets via GetCountsOrThrow.
Capabilities.Float16VectorType is never read anywhere in src/, so a SqlVector<Half> on a v1-negotiated connection fails server-side rather than client-side. (Float32VectorType was already dead in main; this adds a second.)
The negotiation theory's 0x3 case doesn't test what its comment claims — the simulated server caps the ack itself, so the client's own ceiling check at SqlConnectionInternal.cs:1660 stays untested.
Docs say narrowing "fails for values outside its range"; the code saturates to ±Infinity and leaves it to the server.
MetaData is dereferenced without a null check in CreateSourceColumnMetadata.
Things I checked and found correct
Worth recording, since they're the parts most likely to be wrong in a change like this:
Subnormal widening, signed zero, overflow to infinity, the flush-to-zero boundary (2⁻²⁵ ties to even → 0, 2⁻²⁶ flushes), binary32 subnormal inputs, and the rounding carry into the exponent are all correct. Compiling Manual* on every TFM so the netfx path is under test on .NET too is a good arrangement.
Bulk copy null handling is correct as-is — GetValueFromSourceRow returns DBNull.Value with isNull = true and ConvertValue returns before ConvertVectorToBaseType. Note a5de733c3 is test-only; it documents behaviour that already worked rather than fixing anything.
The raw-payload path can't be taken by DataTable, DataRow[], or non-SqlClient DbDataReader sources (they keep ValueMethod.GetValue), and reordering column mappings are safe because sourceOrdinal is the mapped ordinal.
No unacknowledged breaking changes beyond the three listed: GetDataTypeName is unchanged, GetFieldType/GetProviderSpecificFieldType both route through the single new GetVectorFieldType so they stay consistent, SqlMetaDataFactory.DataTypes only adds a row (gated on MinimumVersionKey), the float32 declaration is deliberately unchanged, and SqlDbColumn's new indexer falls through to base[property].
No shared mutable state: Float16Converter is stateless and SqlVector<T> is a readonly struct with no static caches.
On testing
Taking as given that CI has no float16-capable server and no Azure SQL DB connectivity, so the manual suite is the only gate that will ever run — I looked at whether it is complete enough for a lab run rather than whether CI covers it. It is substantial: 11 behaviour tests plus the inherited NativeVectorTestsBase matrix, and the sample data is well chosen (Half.MaxValue, Half.Epsilon, -0.0f, exactly-representable eighths). Gaps I'd close:
.NET Framework gets none of the NativeVectorTestsBase matrix — NativeVectorFloat16Tests.cs is entirely #if NET. That is where the hand-rolled Manual* codec is the production path.
No async coverage for the new representation on any framework, and none at all for float16 on netfx.
DataTestUtility.CheckVectorFloat16Supported fails open through the code under test. It reads the probe vector with GetString + JsonSerializer.Deserialize and catches JsonException → false. A driver regression that produces malformed JSON silently skips the whole float16 suite green. Since the manual run is the only gate, that is the wrong failure mode. (Outside this diff, so no inline comment — but worth fixing alongside. It also leaves PREVIEW_FEATURES = ON on the shared test database as a side effect.)
Two range tests assert only that someSqlException was thrown; they'd pass on an unrelated failure, and they can't distinguish "client rejects" from "client saturates and server rejects" — which is exactly the ambiguity in the doc wording above.
Nothing covers the blocking issue: reading a float32 column as SqlVector<Half>, for a NULL and a non-NULL row.
Bulk copy with a dimension-count mismatch between source and destination is uncovered. Mitigating: float32→float32 now routes through the new raw-payload path too and is covered by the existing NativeVectorFloat32Tests, which runs against any vector-capable server — so regression risk to shipped functionality is covered.
Minor
SqlVector<T>.ToString() changing for existing SqlVector<float> callers is justified and correctly surfaced in the ref assembly and docs, but it is unrelated to float16 — it wants its own release-note entry as a behavioural break, not just an API-list line. Related: GetString() uses JsonSerializer.Serialize, which throws on NaN/Infinity by default, and on .NET Framework that is now the default GetValue() path.
ConvertPayloadElementType is internal static on SqlVector<T> but never uses T, so callers write SqlVector<float>.ConvertPayloadElementType(...), which reads as though it returns a float32 result.
On .NET Framework the reader surfaces a float16 column as string while an output parameter surfaces it as SqlVector<float>. Self-consistent, but worth documenting.
Preprocessor directives in the new code are indented to the surrounding block; the dominant style in these files is column 0.
ConnectionCapabilities.cs:179 says vectors were "introduced in SQL Server 2022" — pre-existing, but the new Float16VectorType doc sits right beside it.
Worth confirming the doc/samples build resolves the locally-packed driver: SqlVectorFloat16Example.cs references SqlVector<Half> and GetSqlVector<Half>, which exist in no released package.
Review assisted by GitHub Copilot; findings verified against the code at a5de733c3.
@apoorvdeshmukh and @cheenamalhotra, I think we will need to call this API a known limitation for preventing backward migration from .Net Runtime to NetFx.
BTW, I saw changes to ref assembly. Should I expect two copies of changes, one for netcore and another for NetFx? I am curious about how the APIs will show up in contract assemblies targeting 2 different frameworks.
Correctness
- Reject a narrowing read consistently for null and populated rows. The
element type is now checked before the null check in GetSqlVector<T>, so
reading a float32 column as SqlVector<Half> fails for every row rather
than succeeding for the null ones.
- Validate the vector header's magic number and version on the widening and
payload conversion paths, which previously checked only the length.
- Quieten a signalling NaN when widening to single precision, and preserve a
NaN's sign and payload in both directions, so the hand written codec and
System.Half agree on all 65,536 bit patterns.
- Read the redirected scale when converting a bulk copy value, so an
encrypted column uses its base type rather than the wrapping metadata's.
- Guard against a null MetaData when deciding whether a bulk copy source can
supply a raw vector payload.
Behaviour
- Report a value which cannot be narrowed to float16 during a bulk copy as an
OverflowException, rather than saturating it to an infinity and letting the
server reject the result as a malformed vector.
- Remove the unused Float32VectorType and Float16VectorType capability
properties. The negotiated version is still recorded in VectorVersion. A
client side guard was considered in their place, but the server already
reports an unrecognised base type clearly, so the guard would only have
replaced a good error with a worse one, and would have made float16 fail
differently from float32 for the same cause.
- Return the JSON rendering of a float16 vector as a SqlString from the
provider specific accessors, so that every provider specific value remains
a type from System.Data.SqlTypes. GetValue continues to return a string.
Tests
- Run the whole native vector matrix against a float16 column through the
single precision representation, which covers .NET Framework, where the
hand written codec is the production path rather than a test double.
- Assert NaN bitwise rather than skipping it, which is what allowed the
codec divergence above to go unnoticed.
- Cover reading a float32 column as a narrower vector, for null and populated
rows, synchronously and asynchronously.
- Exercise the client's own feature extension version ceiling, by letting the
simulated server acknowledge a version regardless of what the client
requested. The existing case only proved the harness capped the version.
- Assert the server's error number for an out of range value rather than
accepting any SqlException.
- Show the column metadata driving a read for a caller which does not know
the schema in advance.
- Remove the SqlVector<T>.ToString() override added earlier in this branch.
It changed the rendering of the already shipped SqlVector<float> as well, and
the reader already exposes the JSON form through GetString and
GetFieldValue<string>. The internal GetString is unchanged.
- Move Float16Converter into the Microsoft.Data.Common namespace, matching the
folder it lives in and its neighbours there.
Docs
- State that a bulk copy reports an out of range narrowing itself, and
correct the SQL Server version for the float32 base type.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: e17ed782-c576-4cb7-9b4b-7ad286d7a7d0
The reason will be displayed to describe this comment to others. Learn more.
Copilot encountered an error and was unable to review this pull request. You can try again by re-requesting a review.
Note
This error may be related to your runner configuration. You can now configure runners for Copilot code review separately from Copilot cloud agent by creating a copilot-code-review.yml file with your setup steps. Read the docs for details.
The reference test steps through the single precision space with a prime
stride, so it samples about 4.1 million of the 4,294,967,296 patterns rather
than visiting them all. Say so, and record why: a full sweep takes about ten
seconds on 32 cores and proportionally longer on a smaller agent, which is not
worth repeating on every CI leg for a codec that does not change. It was run
once and reported no divergence, and setting the stride to 1 reproduces it.
Also document the connection string keyword synonym test, which covers every
accepted spelling through both the builder and a connection.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: e17ed782-c576-4cb7-9b4b-7ad286d7a7d0
[P2] Keep widened server values from becoming oversized sendable vectors. A valid vector(1999..3996, float16) is widened here into SqlVector<float> with _size = 8 + 4 * length, which exceeds the 8000-byte float32 limit even though the constructor bypasses validation. Because this result is public and can be assigned to a SqlParameter, the parameter path later declares/sends it as a float32 vector and fails for these otherwise readable columns. Preserve a representation that can be converted back to the source/destination base type, or explicitly prevent oversized widened read values from being reused as parameters.
The reason will be displayed to describe this comment to others. Learn more.
Float16 compatibility review (head 92d712a)
Checked against the design notes: default V1, opt-in V2, native Half on .NET, JSON plus SqlVector<float> widening on .NET Framework, minimal client conversion, and room for future base types. The basic direction matches. The gaps below are about mixed setups: different negotiated levels, frameworks, and server builds.
Must address
Native float16 is sent on V1/Off connections. See the inline comment on SqlCommand.cs.
SqlVector<float> in a DataTable cannot be bulk copied into a float16 column. This is the only strongly typed write form on .NET Framework. See the inline comment on SqlBulkCopy.cs.
Writes across base types depend on the server. The public SQL docs say float32↔float16 conversion is blocked. See the inline comment on SqlVector.xml.
Still open: V2 source → V1/Off (or varchar) bulk copy destination fails on .NET.SqlVector<Half> reaches Convert.ChangeType instead of becoming JSON. Reported in #4501 (review) and not fixed. The same failure happens with a native float16 source and a varchar/nvarchar/json destination column. It also runs against the design goal of moving JSON through bulk copy transparently. A small fix in the CoerceValue string branch (ISqlVector → JSON) would cover all of these.
Still open: wide widened values (1,999–3,996 elements) can be reused as parameters. They then send a float32 payload larger than 8,000 bytes. Reported in #4501 (review). Please fail early on the client with a clear message, or document that JSON is the write-back path.
Show that the float16 suites really ran. The ADO build 179897 for this PR shows PartiallySucceeded. CheckVectorFloat16Supported also returns false on a JsonException, so a driver bug that breaks JSON output would quietly skip every float16 suite. It also turns on PREVIEW_FEATURES for the shared test database. Raised in #4501 (review) and still present. Please confirm with test results that float16 tests ran (not skipped) on at least one net462 leg and one .NET leg.
Good to have
On .NET Framework, a DataAdapter/SqlCommandBuilder update of a V2 float16 column fails (inline on SqlDataReader.cs).
Bulk copy rounds JSON twice (decimal → float32 → float16), so results may differ from server-side JSON conversion (inline on SqlBulkCopy.cs).
Vector metadata is hard to discover, and it only works on negotiated connections (inline on SqlDbColumn.cs).
Document that cross-base-type bulk copy of reader payloads behaves differently by framework: rejected on .NET, converted on .NET Framework. Today SqlVector.xml describes only the .NET behavior.
Low priority
The ConnectionString keyword table in doc/snippets/Microsoft.Data.SqlClient/SqlConnection.xml does not list Vector Type Support.
On .NET Framework, a float16 output parameter returns SqlVector<float> from Value but SqlString from SqlValue (inline).
JSON → float32 bulk copy lost the client-side 1,998-element check (inline).
The sample states the wrong minimum version (inline).
Question: on .NET Framework, will the JSON text a float16 column returns change when an app moves from V1 to V2? Under V1 the server writes the text; under V2 the client builds it. If the text changes, add a note in the docs.
Scope: I read the source only. I did not run builds or tests. All 44 inline threads are already resolved, and I checked the earlier fixes against the current code instead of relying on that status.
Nothing checked the negotiated version before a float16 value was declared and
written. Measured against SQL Server vNext CTP 1.0: a V1 connection sent the
value and the server stored it, even though the same connection reads that
column back as a JSON string, and an Off connection was told its protocol
stream was malformed (error 8070). Both are worse than naming the keyword,
which is what the JDBC and ODBC drivers do. Checked once in
VectorTypeSupportUtilities so the parameter and bulk copy paths cannot
disagree.
Convert any in-memory value to the destination's base type during a bulk copy,
not only a textual one. A SqlVector<T> carries the base type of its element
type, which the caller chose rather than the column, so a SqlVector<float> from
a DataTable reached a float16 column at the wrong width and the server rejected
it. The test matrix had been narrowed to hide this; it now runs both source
modes for every representation. A payload read from another vector column is
still left alone, so a cross base type copy is still reported by the server.
Say that conversion between base types depends on the server. Some builds block
it and report error 42238, which leaves a JSON string as the only portable way
to write a column whose base type differs from the value's, and the only way at
all on .NET Framework. Tests which depend on the conversion now skip rather
than fail where it is blocked.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: e17ed782-c576-4cb7-9b4b-7ad286d7a7d0
The JSON intermediate is built without the element count limit so that a wide
float16 column can be loaded through it, which also removed the check for a
float32 destination: an oversized array reached the server and came back as a
column length error. The limit depends on the element width, so it is now
applied after the payload has been rewritten to the destination's base type.
Cover the negotiation check with tests, and record what a connection below the
negotiated version actually does during a bulk copy: such a connection is told
a float16 column is a varchar(max), so the value travels as text and the
server converts it, which is what makes the default level usable for loading
float16 data. The check on that path guards the invariant rather than a
reachable case.
Add a rounding parity test. Text bound for a float16 column is parsed to
float32 and then narrowed, so it is rounded twice, while an INSERT leaves the
text to the server. Measured at and around binary16 ties, including the
smallest subnormal: both paths agree.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: e17ed782-c576-4cb7-9b4b-7ad286d7a7d0
The newly added test methods in this class do not have behavior-focused XML <summary> documentation, despite the repository testing requirement for every [Fact]/[Theory] method. Add summaries to these tests (and the other undocumented methods in this file) describing the contract and why each scenario matters.
The declaration and the binary metadata for a vector parameter were read by
casting the raw value to ISqlVector, so a JSON string paired with
SqlDbType.Vector threw InvalidCastException. That pairing is what a
DbDataAdapter update of a float16 column produces on .NET Framework: the column
has no System.Half to be surfaced as, so it is read as a string, and
SqlCommandBuilder takes SqlDbType.Vector from the column's provider type.
Both forms are now described through one helper, so the declaration and the
payload agree whichever was used. Nothing accepted a string here before, so
this only widens what works.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: e17ed782-c576-4cb7-9b4b-7ad286d7a7d0
VectorBaseType and VectorDimensions are null for a vector column the server
returned as a varchar(max), which is every float16 column at v1 and every
vector column at off. An application on the default therefore cannot tell a
float16 column from text through them, so point at sys.columns.vector_base_type
for that case, and cover it with a test.
Record that a float16 output parameter cannot be declared from .NET Framework,
since both value forms available there declare float32. Measured: Value and
SqlValue agree, both returning a widened SqlVector<float>. The branch is kept
for a server which returns the column's own base type regardless.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: e17ed782-c576-4cb7-9b4b-7ad286d7a7d0
[P3] Add XML summaries to the new test methods in this class. The repository testing guidance requires behavior-focused summaries for every test method (.github/instructions/testing.instructions.md:181-190), but these methods rely only on inline comments. Document the contract and regression purpose for each method so the new float16 behavior suite follows the test documentation convention.
[P2] Skip this adapter test when base-type conversion is unavailable. The adapter supplies a string value with SqlDbType.Vector, and SqlParameter.GetVectorProperties parses that string into a float32 binary payload; it is not sent as textual data. Updating a float16 column therefore requires server float32-to-float16 conversion, which DataTestUtility.IsSqlVectorBaseTypeConversionSupported explicitly reports may be unavailable, so this test will fail on supported server builds that reject error 42238.
Building a vector from a JSON string fixed its base type to float32, so a
column of another base type then needed the server to convert, which not every
build does. The string is now sent as a varchar(max) for the server to parse
into whatever base type the column has, which works everywhere and is the only
form available to a .NET Framework caller writing a float16 column.
Covered by writing 3000 elements to a float16 column, which no float32 payload
could carry: a vector payload is capped at the size of a TDS packet, so 1998
float32 elements. It only succeeds if the value travelled as text.
Seed the float16 rows of read tests from a literal rather than a
SqlVector<float> parameter, so that setting up a test does not depend on the
server converting between base types when the test is not about that.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: e17ed782-c576-4cb7-9b4b-7ad286d7a7d0
Inserting a test between a summary and its method left two summary blocks
together, detaching the one for DrivesReadPathForACallerWhichDoesNotKnowTheSchema.
Put each back above the method it describes.
The testing instructions ask for a behaviour-focused summary on every test
method, so promote the leading comments in VectorFloat16BehaviourTests to
summaries as well; those were written before that requirement was raised and
were the only ones in the vector suites still relying on inline comments.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: e17ed782-c576-4cb7-9b4b-7ad286d7a7d0
[P3] Use ASCII punctuation in source comments. The repository coding style requires non-ASCII source characters to be escaped, but these new em dashes are literal characters in a .cs file; replace them with ASCII hyphens (or Unicode escapes).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Area\VectorUse this for issues that are targeted for the Vector feature in the driver.
7 participants
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds support for SQL Server's
vector(N, float16)base type, opted into by a new connection string keyword.System.Halfexists on .NET but not on .NET Framework, so the two differ by necessity:float16column is read and written asSqlVector<Half>, exchanged exactly as the server sent it.SqlVector<float>, which widens losslessly. Writes go back as a JSON string or aSqlVector<float>.Opting in
Vector Type Support(also accepted asVectorTypeSupport) takesoff,v1orv2, and defaults tov1— the same name, values and default as the JDBC driver'svectorTypeSupport. Upgrading the driver therefore does not change the representation an existing application receives;float16arrives in binary form only when the application asks forv2.A base type can only be used once the connection has negotiated the version covering it. The driver rejects a value above that version itself, naming the keyword to set, rather than leaving it to the server — which was measured to accept it silently at
v1and to report a malformed protocol stream atoff. The acknowledgement is likewise refused if it exceeds the version requested.Conversion between base types
Whether SQL Server converts between
float32andfloat16depends on the server: some builds block it and report error 42238. Verified as working against SQL Server vNext CTP 1.0 (18.0.258.0). A JSON string is accepted by every server, so it is the portable way to write a column whose base type differs from the value's — and on .NET Framework the only way where conversion is blocked. Tests which depend on the conversion skip rather than fail.Bulk copy
The
INSERT BULKstatement states the destination column's base type, and the server then requires a binary payload of exactly that width and performs no conversion within the data stream. Measured at every negotiated version: a text declaration for a vector column is refused withInvalid column type from bcp client, and text sent under a vector declaration is refused with a length mismatch.So any in-memory value — a JSON string, or a
SqlVector<T>whose element type differs from the column's — is converted to the destination's base type by the driver, with a range check, which does not depend on the server. A payload read from another vector column keeps its own base type, so copying between columns of different base types is reported by the server rather than silently narrowed. Both match the JDBC driver.A connection below
v2is told afloat16column is avarchar(max), so bulk copy into one still works at the default level: the value travels as text and the server converts it.Public API
One addition: the
SqlVectorTypeSupportenum andSqlConnectionStringBuilder.VectorTypeSupport.SqlVector<T>is unchanged.Tests
float16added to the existing generic native vector suite, run over av2connectionSqlVector<float>on every framework, which is the .NET Framework representation and where the hand written binary16 codec is the production pathSystem.Halfacross all 65,536 binary16 patterns, and across the binary32 space by sampling (a full sweep of all 4,294,967,296 patterns was run once during development and reported no divergence, but is too slow to repeat on every CI leg)INSERTDbDataAdapterround trip of afloat16column on .NET FrameworkChecklist